Back

Scientific Data

Springer Science and Business Media LLC

Preprints posted in the last 90 days, ranked by how well they match Scientific Data's content profile, based on 209 papers previously published here. The average preprint has a 0.15% match score for this journal, so anything above that is already an above-average fit.

1
A large-scale crowd-sourced annotated acoustic dataset of Indian fauna

Ramesh, V.; Singh, S.; Pop, P.; Choksi, P.; Singh, P.; Khanwilkar, S.; Teotia, S.; Burli, P.; Devarajan, K.; A, A.; A Nakhwa, A.; Abdus Shakur, M.; Baishya, R.; Bhagwat, N.; Biniwale, S.; Bora, C.; C S, S.; Chakraborty, N.; D'Souza, S.; D'Souza, E.; Vaishnav, R. D.; Deshpande, K.; Dhanda, A.; G, A.; Ghosh, A.; Goswami, R.; K N, A.; K P, N.; K Rajaraman, B.; K V, G.; Kannan, V.; Karthick, V.; Kotian, M.; Kumar, H.; Kurian, P.; Madhavan, M.; Meena, K.; Mohammad Maslehuddin, A.; Mourya, P.; Mudke, M.; R J, P.; R S Jha, R.; Ramesh, K.; Sailas, S. S.; Sangwan, T.; Mahesh, S.; Satish, R.; Shankar, A

2026-07-21 ecology 10.64898/2026.07.20.739496 medRxiv
Top 0.1%
75.2%
Show abstract

Global rates of biodiversity loss warrant conservation action and monitoring at large geographic scales. Conservation technologies such as acoustic monitoring in conjunction with deep learning now enable us to monitor wildlife simultaneously across space and time. However, for a significant proportion of biodiversity in tropical regions, we cannot yet rely on automated recognition approaches because we lack acoustic templates to robustly train deep learning algorithms. In this paper, we relied on a novel participatory approach, enlisting researchers, conservation practitioners, and nature enthusiasts to create a unique crowd-sourced, open-access dataset of acoustic annotations across taxonomic groups for biodiversity in India. Our dataset comprises 3311 minutes of strongly labelled data (bounding boxes or annotations for a species vocalization) and 2504 minutes of weakly labelled data (indicating the presence of a species within an audio file but lacking bounding boxes) for 518 species across India, spanning 25 of 36 states and union territories. We present metadata and code for data processing and highlight the strengths of a participatory approach to biodiversity monitoring.

2
Hybrid transcriptome assembly and annotation of Japanese macaque prefrontal cortex

Chatzipli, A.; Voshall, A.; Viswanadham, V.; Weiss, A. R.; Liguore, W. A.; McBride, J. L.; Sherman, L. S.; Lee, E. A.; Yu, T. W.

2026-08-19 neuroscience 10.64898/2026.08.10.743719 medRxiv
Top 0.1%
56.2%
Show abstract

Japanese macaque (Macaca fuscata) is used in biomedical and neurobiology research, yet transcriptomic resources for the brain are limited. We present a hybrid RNA sequencing dataset and a prefrontal cortex transcriptome assembly from two healthy 6-year-old animals. Short-read Illumina ({approx}70 million paired-end reads per sample) and long-read Oxford Nanopore direct RNA sequencing ({approx}2.5 million reads per sample) were combined. Reads were quality controlled, aligned to the macFus_1.0 reference genome, and assembled with StringTie2. Transcripts were annotated using Trinotate and eggNOG-mapper, and open reading frames were predicted with TransDecoder. The released data package includes raw reads (NCBI SRA BioProject PRJNA1295993), transcript sequences and structural annotation files, predicted coding sequences and proteins, functional annotation tables, and transcript abundance estimates (TPM). Technical validation includes read-level QC and protein-level comparisons to expressed gene sets from human, rhesus macaque and chimpanzee prefrontal cortex. These resources enable reuse for transcript-level expression studies, isoform characterization and comparative primate neurogenomics.

3
A versioned, analysis-ready archive of United States State Cancer Profiles county- and state-level estimates

Davis, S.

2026-08-27 health informatics 10.64898/2026.08.24.26361254 medRxiv
Top 0.1%
46.5%
Show abstract

State Cancer Profiles (statecancerprofiles.cancer.gov), maintained by the National Cancer Institute with the Centers for Disease Control and Prevention, is a widely used source of county- and state-level cancer statistics in the United States, used for cancer-center catchment-area surveillance and for geographic studies of cancer burden, screening, and access to care. The site offers no API, no bulk download, and no archive of prior estimates: its sole machine-readable export returns one statistical stratum per HTTP request, and when the underlying data are updated the previous estimates are overwritten and become unrecoverable. This resource provides the complete national county- and state-level extract of all four State Cancer Profiles data topics (incidence, mortality, screening and risk factors, and demographics) as typed, analysis-ready files with the stratifying dimensions as columns, published under pinned, citable version DOIs on Zenodo (concept DOI 10.5281/zenodo.11098814). One version DOI is minted per distinct upstream data vintage, the set of values the site served between successive replacements. Three vintages have been captured to date; at each observed vintage boundary roughly 97% of estimate values changed, so which vintage an analysis draws on affects its results. From the 2026-08-24 release forward, cells that the upstream site suppresses are retained as typed nulls with an explicit suppression-reason column. Capture has been automated on an approximately monthly cadence since February 2025, and each future upstream revision will be preserved as a new vintage.

4
A sex-balanced longitudinal developmental task-free fMRI rat dataset.

Mc Loone, D. P.; Breen, A.; McParland, C.; Harkin, A.; Kelly, C.

2026-06-09 neuroscience 10.64898/2026.06.04.730202 medRxiv
Top 0.1%
45.6%
Show abstract

Here we present a preclinical longitudinal developmental dataset spanning pre-puberty (juvenile) to early adulthood in Wistar rats. Thirty-six rats (18 female) were scanned during five sessions at approximately postnatal day 28, 35, 49, 70 and 91. A standardised consensus protocol was used for animal sedation and MRI data acquisition (structural MRI and resting-state fMRI). The dataset also includes daily body weight measurements, day of puberty onset, and daily cytological smear images for oestrous tracking. This openly available dataset addresses the need for sex-balanced developmental preclinical datasets spanning puberty and enables researchers to investigate longitudinal brain functional developmental trajectories and Sex As a Biological Variable (SABV).

5
Brainana: an end-to-end preprocessing framework for macaque neuroimaging

Liu, X.; Zhang, Y.; Yin, Z.; Zhen, Z.; Arcaro, M. J.

2026-06-08 neuroscience 10.64898/2026.06.03.729972 medRxiv
Top 0.1%
34.8%
Show abstract

Macaque MRI bridges non-invasive systems neuroscience with cellular and circuit-level mechanisms, but preprocessing remains fragmented across tools that are difficult to integrate, adapt to non-human primate acquisitions, and deploy reproducibly. We present Brainana, an automated, BIDS-compatible preprocessing framework for macaque neuroimaging. Brainana integrates structural and functional preprocessing, cortical surface reconstruction, quality control, transform tracking, and atlas projection within a containerized package, with cloud access for users without local compute. It incorporates macaque-trained deep learning models for brain extraction and tissue segmentation, conformation to standardize variable acquisitions, and surface reconstruction optimizations for macaque neuroanatomy. Across 23 imaging sites, Brainana processed data spanning heterogeneous scanners, protocols, species, and resolutions, yielding accurate anatomical correspondence across 130 monkeys, reliable native-space cortical surfaces, localized task-evoked activations, and reproducible brain-wide resting-state correlation structure. Brainana enables reproducible, scalable, and accessible macaque MRI preprocessing that supports cross-study comparison and multimodal integration across spatial scales, from neurons to networks.

6
Iterative co-creation of harmonized human and non-human primate cellular and structural ontologies and 3D common coordinate frameworks for the basal ganglia

Ding, S.-L.; Bhandiwad, A.; Rosen, B.; Seeman, S. C.; Long, B.; Johansen, N. J.; Bayindir, U.; Facer, B.; Fu, Y.; Halimi, Y.; Hou, Y.; Hu, D.; Huang, M.; Ikeda, T.; Kalmbach, B.; Kruse, L.; Lesnar, P.; Liu, X.-P.; Luo, Z.; Ray, P.; Royall, J. J.; Schmitz, M. T.; Uematsu, A.; Vezoli, J.; Yazdani, F.; Bakken, T. E.; Freiwald, W.; Hayashi, T.; Hodge, R. D.; Kennedy, H.; Mollenkopf, T.; Ng, L.; Osumi-Sutherland, D.; Thompson, C. L.; Hawrylycz, M.; Glasser, M. F.; Van Essen, D. C.; Zeng, H.; Lein, E. S.

2026-08-04 neuroscience 10.64898/2026.07.30.741796 medRxiv
Top 0.1%
34.8%
Show abstract

A major goal of the BRAIN Initiative Cell Atlas Network (BICAN) is to create a suite of foundational reference cell atlases and associated standards for human and non-human primate brains. Central to this goal is the creation of cross-species harmonized cellular taxonomies and structural parcellations with formal ontologies that can be mapped into 3D reference frameworks bridging neuroimaging and cellular and histological resolutions. We describe here an iterative approach, focused initially on the basal ganglia, to co-create structural and cellular ontologies in human, macaque and marmoset brains, including a Harmonized Ontology of Mammalian Brain Anatomy (HOMBA), and to map and refine structural parcellations into neuroimaging-based common coordinate frameworks. These references provide the framework for documenting and mapping all experimental sampling in BICAN, allowing analyses of cellular and molecular variation as a function of topographic position, and enabling comparisons of cellular, molecular and neuroimaging-based functional variation within and between primate species. HighlightsO_LIA hierarchical Harmonized Ontology of Mammalian Brain Anatomy (HOMBA) covering 2348 structures C_LIO_LIHOMBA-annotated 3D common coordinate frameworks (CCFs) of the basal ganglia across species C_LIO_LIHistologically informed 3D parcellation/atlas of 280 human subcortical structures indexed by HOMBA C_LIO_LIMapping and integration of structural, cellular and functional data with HOMBA and CCFs C_LI

7
Whole-genome sequencing data of a diverse grapevine germplasm collection maintained in Bordeaux, France

de Miguel, M.; Lafargue, M.; Saez-Laguna, E.; Tran, J.; Girollet, N.; Bert, P.-F.; Wang, Y.; Liang, Z.; Guillaumie, S.; Dai, Z.; Ollat, N.

2026-08-05 genomics 10.64898/2026.07.31.742002 medRxiv
Top 0.1%
34.3%
Show abstract

Grapevine (Vitis vinifera) is one of the worlds most economically important fruit crops and a model species for perennial fruit tree genetics and genomics. The extensive genetic diversity found in cultivated and wild Vitis species provides a valuable resource for studies of domestication, adaptation, trait evolution, and breeding. This article presents a standardized whole-genome variant dataset comprising 547 grapevine accessions maintained in the INRAE Bordeaux grapevine germplasm collection, including 397 domesticated V. vinifera cultivars and 150 wild Vitis accessions. Whole-genome sequencing data were generated at a target sequencing depth of approximately 20x, and sequence variants were identified using a standardized Genome Analysis Toolkit (GATK) workflow against the reference genome PN40024v4 (40X). Variant discovery was carried out simultaneously across the complete sample set to ensure consistent genotype calling. All accessions were sequenced using the same technology and processed using the same reference genome, sequence alignment, variant-calling, and filtering workflow to produce a standardized variant dataset comprising ca. 9.1M SNPs and 0.77M INDELs. The resulting VCF files provide a harmonized genomic resource that can be readily reused for studies of grapevine genetics, germplasm characterization, population genomics, comparative genomics, genome-wide association studies, and the development and benchmarking of bioinformatic methods. SPECIFICATIONS TABLE O_TBL View this table: org.highwire.dtl.DTLVardef@490449org.highwire.dtl.DTLVardef@1b884fborg.highwire.dtl.DTLVardef@12280bforg.highwire.dtl.DTLVardef@32bcceorg.highwire.dtl.DTLVardef@1099221_HPS_FORMAT_FIGEXP M_TBL C_TBL VALUE OF THE DATAO_LIThis dataset provides whole-genome raw sequencing for 118 wild Vitis accessions originating from North America and Asia and a sequencing-derived variant data for these accessions and 429 Vitis vinifera cultivars and wild accessions from the INRAE Bordeaux germplasm collection, previously published by Dong et al. 2023[1]. The dataset captures genetic variation across a total of 547 grapevine accessions, including domesticated and wild grapevine germplasm using a common variant-calling pipeline, facilitating direct comparisons among accessions. C_LIO_LIThe inclusion of wild Vitis species together with cultivated grapevine accessions provides a resource for studies of grapevine diversity, domestication, phylogenetic relationships, and comparative genomics. The dataset enables the investigation of genetic variation across multiple Vitis species using a standardized set of genomic variants. C_LIO_LIThe variant call format (VCF) file can be reused for population genetics, phylogenetic analyses, genetic diversity assessments, introgression analyses, and the identification of genomic regions of interest. The dataset is compatible with widely used bioinformatics software and can be integrated with other publicly available grapevine genomic resources. C_LIO_LIThis dataset constitutes a genomic resource for grapevine breeding and conservation research. Researchers can use these data to identify genetic diversity in wild relatives, compare allelic variation between cultivated and wild germplasm, investigate candidate loci associated with traits of interest, and support the management and characterization of grapevine germplasm collections. C_LI

8
HyenaSET: Hyena Sound Event Transcripts and benchmark animal2vec performance for parsing animal communication

Woerner, J. M.; Angonin, C.; Gersick, A. S.; Holekamp, K. E.; Jensen, F. H.; Johnson, M. P.; Onsare, M. H. M.; Pioon, M. O.; Schäfer-Zimmermann, J.; Strandburg-Peshkin, A.; Strauss, E. D.

2026-06-17 animal behavior and cognition 10.64898/2026.06.14.732108 medRxiv
Top 0.1%
34.2%
Show abstract

Here, we present HyenaSET, a large ([~]1640 hours) bioacoustic dataset derived from collar-mounted audio recorders deployed on 19 spotted hyenas in the Maasai Mara National Reserve, Kenya. Within this dataset, 243 hours have been strongly labeled by identifying the onset and offset of all vocalizations as well as their types, and the labels have been validated by experts on hyena vocalizations and behavior. Within the strongly labeled data, the total amount of time hyenas were vocalizing was 9.5 hours (3.9%). Furthermore, each vocalization has been manually identified as "focal" (emitted by the hyena wearing the collar) or "non-focal" (emitted by a nearby conspecific), making use of information from collar-mounted accelerometers that picked up vibrations of the animals throat when it produced vocalizations. In addition to the labeled data, we also provide a large corpus of unlabeled data from the same recordings, which can be used for un- or self-supervised machine learning tasks. To ensure reproducibility of this dataset as a benchmark in machine learning studies, we present it alongside five stratified cross-validation train/test splits to enable accurate comparisons, and we also provide a train/test split in which specific individuals are left out of the training set to assess generalizability across individuals. Finally, as a performance benchmark, we present baseline results for this dataset using animal2vec, a recently developed transformer-based model optimized for bioacoustic data.

9
Bridging Biomedical Atlas Ecosystem: Cross-Atlas Alignment And Scalable Tissue Specimen Registration

Jain, Y.; Desai, B.; Qaurooni, D.; Bhavsar, A.; Kienle, P.; Pouch, A. M.; ONeill, K.; Apte, S.; Herr, B. W.; Fisher, S. A.; Börner, K.

2026-08-22 bioinformatics 10.64898/2026.08.13.744704 medRxiv
Top 0.1%
31.7%
Show abstract

Over the last five years, over 13,000 tissue datasets with 200+ million cells from 20 consortia have been spatially registered into the Human Reference Atlas (HRA) common coordinate framework (CCF). The shared 3D spatial and semantic reference system enables exploration of datasets in the context of all other data across organs, assay types, and spatial scales. However, manual registration of individual samples remains resource intensive, posing feasibility challenges exacerbated by the proliferation of samples, assays, and atlasing efforts. This paper presents two approaches to scale up HRA construction: (1) projecting data across biomedical reference atlas systems and (2) using millitomes to bulk register tissue blocks into a reference organ. Both methods use the AutoMated Alignment and Projection (AMAP) pipeline to align 3D mesh models using point cloud registration. We demonstrate the evolving HRA-aligned atlas ecosystem for 6 models from the SPARC Program (heart), Gut Cell Atlas (large intestine), 500-subject consensus kidneys, and the Julich Brain Atlas. Additionally, we used AMAP to project 7 millitome models across 5 organs onto the HRA ecosystem, integrating 300+ tissue extraction sites. AMAP enables scalable tissue registration of data across atlas ecosystems enabling the construction of detailed reference maps of the human body.

10
A behaviourally normed database of 1,377 natural sounds for auditory cognition and neuroscience

Plegat, M.; Araujo Vitoria, M.; Marinato, G.; Tita, B.; van der Lans, C.; Pijfers, M.; Esposito, M.; Bertovic, M.-S.; Formisano, E.; Giordano, B. L.

2026-08-28 neuroscience 10.64898/2026.08.25.746933 medRxiv
Top 0.1%
31.7%
Show abstract

Natural-sound research requires stimulus sets that combine acoustic standardization with detailed behavioural characterization. We present 1,377 two-second sounds representing 240 expert-defined source--action classes. We call this database "MaMa Sounds", as it resulted from the collaborative effort of two academic teams in Maastricht and Marseille. The sounds were manually curated, segmented, sampled at 16 kHz, and labelled with a noun identifying the source and a verb identifying the action. We release deidentified trial-level identification and familiarity data together with multiple per-sound norms (e.g., identification accuracy, confidence and agreement; familiarity), along with overall norms derived with principal component analysis. Noun, verb, and joint noun--verb norms are provided as direct means and medians with the number of contributing observations. This battery preserves process-specific information, while two principal-component scores provide compact overall behavioural-identifiability measures derived from response ease, semantic correspondence, agreement, and familiarity. The repository also contains deterministic response-cleaning code, participant and reference Word2Vec representations, and code reproducing the public sound-level tables. The resource supports stimulus selection, matching, and continuous modelling in auditory cognition and neuroscience.

11
A Modular, AI-assisted Digitization Toolkit for Resource-Constrained Herbaria: A Case Study from Zimbabwe

Gatula, L.; Bezrukov, I.; Atemia, J.; Chapano, C.; Gamundani, P. T.; Zimudzi, C.; Chatukuta, P.

2026-07-22 plant biology 10.64898/2026.07.20.739607 medRxiv
Top 0.1%
31.2%
Show abstract

Herbaria serve as invaluable spatio-temporal repositories of plant diversity information. Digitization of herbarium collections enhances the accessibility, discoverability, and long-term preservation of this important plant information, yet financial and infrastructural constraints often prevent herbaria in resource-constrained regions from digitizing their collections. Consequently, critical plant diversity data gaps remain due to underrepresentation of these collections in global biodiversity databases. Here, we describe an AI-assisted modular digitization toolkit specifically designed for herbaria operating under limited funding, developed and refined through our experience digitizing the crop wild relative (CWR) collection of the National Herbarium of Zimbabwe. The toolkit comprises three core components: (1) a portable, cost-effective photostation assembled from commodity parts, (2) a streamlined cascade workflow for systematic digital imaging, and (3) an AI-assisted data management pipeline for image quality control, label transcription, data analysis, and presentation. Compared to manual transcription and legacy optical character recognition approaches, AI-based transcription achieves lower time cost while maintaining high accuracy, and AI-driven data management delivers accessibility and reduced expenditure relative to conventional database infrastructure. The toolkit is designed to allow herbarium staff full autonomy over the digitization procedure, ensuring institutional ownership and the capacity for independent continuation beyond initial project support. By prioritizing affordability, modularity, and simplicity, this toolkit provides a replicable framework that may enable resource-constrained herbaria to locally generate high-quality scientific data for conservation and the sustainable utilization of plant genetic resources.

12
Chromosome-scale genome assembly and annotation of the Vietnamese indica rice cultivar Khang Dan 18

Nguyen, T. Q.; Do, K. H. D.; Vu, T. M.; Hoang, N. V.

2026-08-21 plant biology 10.64898/2026.08.15.742683 medRxiv
Top 0.1%
31.0%
Show abstract

Khang Dan 18 (KD18) is an Oryza sativa L. subsp. indica rice cultivar widely cultivated in northern Vietnam and used as an experimental and breeding background in Vietnamese rice research. Although KD18 has previously been represented in low-depth population resequencing datasets, a contiguous and annotated cultivar-specific genome has not been available. Here, we report a chromosome-scale genome assembly of KD18 generated using Oxford Nanopore long-read and Illumina short-read sequencing. The 395.3-Mb assembly comprises 12 chromosome-scale pseudomolecules containing approximately 95% of the assembled sequence and 99.6% of the predicted protein-coding genes. The assembly showed 97.2% BUSCO completeness, an average Merqury quality value of 46 and a long terminal repeat assembly index of 13.21. A total of 56,546 protein-coding genes representing 71,237 transcripts were predicted, with 99% BUSCO and 98.68% OMArk completeness. These statistics are similar to those of other high-quality genome assemblies that were recently published for different Asian rice cultivars, therefore providing a cultivar-specific genomic resource for research involving KD18 and KD18-derived materials.

13
Multi-area single-cell calcium imaging dataset of the mouse cortex across wakefulness, sleep, and anesthesia

Oomoto, I.; Kiyooka, D.; Oizumi, M.; Murayama, M.

2026-07-24 neuroscience 10.64898/2026.07.20.739676 medRxiv
Top 0.1%
30.0%
Show abstract

We present a reusable dataset of neuronal population activity from multiple cortical areas, recorded from layers 2/3 of the mouse cortex at single-cell resolution during wakefulness, natural sleep (including NREM and REM sleep), and isoflurane anesthesia. Using wide-field two-photon microscopy, we recorded approximately 4,000 to 10,000 neurons per session at 7.65 Hz and provided the spatial coordinates of individual neurons. The repository provides both processed datasets and the corresponding raw imaging movies (TIFF) and electrophysiological recordings (MATLAB format). Processed data are distributed in MATLAB format and include {Delta} F/F fluorescence signals, deconvolved spike estimates, Gaussian-smoothed spike estimates, behavioral state annotations, and metadata. This dataset supports reuse in studies of cortical population dynamics, brain-state-dependent activity, and spatially distributed neuronal organization. It should also be useful for method development, benchmarking, and comparative analyses of large-scale neuronal activity across physiological and pharmacological brain states.

14
Whole genome sequencing and variant discovery in 344 global grasspea (Lathyrus sativus L.) lines

Schreiber, M.; Staples, J.; Emmrich, P. M. F.; Edwards, A.; Martin, C.; Bayer, M.; Raubach, S.; Kilian, B.; Shaw, P. D.

2026-06-09 genomics 10.64898/2026.06.05.730453 medRxiv
Top 0.1%
27.3%
Show abstract

The rapid expansion of genomic data resources for major crops is opening new options for crop improvement, while resources for most underutilised crops lag behind, risking a widening gap in crop improvement. One of these underutilised crops is grasspea (Lathyrus sativus), an ancient crop with modern cultivation centred on South Asia and Ethiopia. We conducted whole genome shotgun sequencing on a global collection of 344 grasspea lines, producing over 152 billion reads. Following variant discovery and filtering we created a single nucleotide polymorphism (SNP) marker set of over 1.5 million SNPs. This is a resource of major significance for this crop which can help unlock its breeding potential through marker development and the identification of genes controlling agronomically important traits.

15
Haplotype-resolved chromosome-level genome assembly of four European white oak species

Magris, G.; Avanzi, C.; Bagnoli, F.; Duvaux, L.; Belmonte, E.; Vendramin, G. G.; Piotti, A.; Pinosio, S.

2026-08-24 genomics 10.64898/2026.08.20.745905 medRxiv
Top 0.1%
27.1%
Show abstract

European white oaks (Quercus section Quercus) are ecologically and economically important forest trees characterized by extensive shared genetic variation and a history of interspecific gene flow. Genomic resources remain uneven across species, limiting comparative analyses and pangenome development. Here, we present haplotype-resolved chromosome-scale genome assemblies and genome annotations for four European white oak species: Quercus robur, Q. petraea, Q. pubescens, and Q. frainetto. The assemblies were generated from PacBio HiFi sequencing data and include both phased haplotypes for each species. Genome sizes range from 779 to 817 Mb and all assemblies are organized into 12 chromosome-scale pseudomolecules with high completeness and contiguity. We additionally provide species-specific repeat annotations, structurally and functionally annotated protein-coding gene sets, and complete organellar genomes. The dataset includes the first reference genomes for Q. pubescens and Q. frainetto, together with newly generated assemblies for Q. robur and Q. petraea produced using a consistent sequencing and analysis workflow. These resources provide a standardized framework for comparative genomics, pangenome construction, genome evolution studies, and investigations of adaptation and introgression across European white oaks.

16
An open-access CT-based 3D anatomical dataset of extant sharks across all major lineages

Yao, S.; Liu, X.; Hou, Y.; Yin, P.; Zhang, X.; Cui, X.; Lu, J.

2026-07-02 evolutionary biology 10.64898/2026.06.29.734410 medRxiv
Top 0.1%
26.8%
Show abstract

Sharks exhibit extraordinary morphological diversity across a wide range of ecological niches, yet large-scale, high-resolution digital datasets of their internal anatomy remain limited. Here we present an open-access 3D shark anatomical repository derived from published X-ray computed tomography (CT) data, featuring manually segmented and systematically annotated models of the chondrocranium, visceral arches, axial skeleton, musculature, and viscera in standard STL format. The dataset comprises 117 individuals, representing 72 species across 25 families and all nine extant shark orders, with 115 full-body reconstructions and two head-only models. This open-access dataset offers a comprehensive resource for comparative anatomy, biomechanical simulations, evolutionary developmental biology and biomimetics research of extant sharks.

17
The Dark Ecology Dataset: Measurements of Aerial Biomass in US Weather Radar from 1995 to 2025

Sheldon, D.; Winner, K.; Deznabi, I.; Bernstein, G.; Bhambhani, P.; Lin, T.-Y.; Desmet, P.; Dokter, A. M.; Horton, K. G.; Nilsson, C.; Van Doren, B. M.; Farnsworth, A.; La Sorte, F. A.; Maji, S.

2026-06-23 ecology 10.64898/2026.06.20.733536 medRxiv
Top 0.1%
22.8%
Show abstract

The US NEXRAD radar network has monitored the aerosphere over the US and its territories continuously since the 1990s and archived nearly 300 million radar volume scans. These data contain a wealth of information about the movements of birds, bats, and insects. Historically, this biological information was difficult to access due to the amount of data and challenges in analyzing it. In the last 15 years, fueled by computational and methodological advances, large-scale aeroecology research has blossomed. However, comprehensive analyses of the NEXRAD archive remain very costly. We collected measurements of biological activity from every volume scan in the NEXRAD archive--nearly 300 million data files total--to assemble a dataset of aerial biomass over the US from 1995 to 2025. The core data are vertical profiles, which summarize biological activity at different heights above the radar station for each volume scan. We also provide time series data products that aggregate vertical profiles to point measurements at radar stations across time. These data products can support a range of aeroecology analyses at significantly reduced effort.

18
WIO-ReefFish: A High-Resolution Dataset for Taxon-Aware Coral Reef Fish Detection in the Western Indian Ocean

Gerard, J.; Branger, L.; Huyghe, F.; Kochzius, M.; Otwoma, L.; Bergacker, S.; op't Roodt, L.; Rumisha, c.; Di Bella, L.

2026-08-20 ecology 10.64898/2026.08.19.745797 medRxiv
Top 0.1%
22.2%
Show abstract

Coral reef fish assemblages are widely used as indicators of ecosystem condition, yet manual annotation of underwater video remains a major bottleneck for scalable biodiversity monitoring. Despite rapid progress in automated detection, ecologically realistic and publicly available datasets remain scarce, particularly for the Western Indian Ocean. Here, we present WIO-ReefFish, a reef fish detection dataset derived from diver-operated line-intercept transects and designed for ecological monitoring under natural survey conditions. WIO-ReefFish comprises 1,000 ultra-high-definition images (3840 $\times$ 2160 pixels) and 6,768 exhaustive bounding-box annotations spanning 24 taxonomic categories, thereby preserving full-frame assemblage structure in complex reef scenes. We also establish a standardized benchmark across nine object detection models under two complementary protocols: class-aware detection and class-agnostic fish localization. Detection performance was consistently higher under the class-agnostic protocol. The best-performing model (RT-DETR) improved from 0.48 mAP50 in the class-aware setting to 0.70 mAP50 when taxonomic constraints were removed, indicating that taxonomic discrimination remains substantially more challenging than fish localisation in reef imagery. Spatially independent evaluation revealed a pronounced generalisation gap, particularly for taxonomic detection, whereas class-agnostic fish localisation remained substantially more robust across transects and countries. Together, these results establish WIO-ReefFish as a realistic benchmark for automated reef fish detection and provide a foundation for more robust computer-vision tools in coral reef biodiversity monitoring. The WIO-ReefFish dataset and associated benchmarking resources are publicly available.

19
Does OMOP CDM Conversion Improve Cross-Country Comparability of Real-World Data? A Benchmark Study in Breast Cancer and Amyotrophic Lateral Sclerosis

Aborageh, M.; Korcinska Handest, M. R.; Bakos, I.; Rajamaki, B.; Silva, C.; Horvath-Puho, E.; Pylkkaenen, L.; Venda, C.; Lentzen, M.; Becker, C.; Fernandes, J.; Paakinaho, A.; Vo, T.; Haenisch, B.; Hartikainen, S.; Tolppanen, A.-M.; Furtado, C.; Froehlich, H.; Ehrenstein, V.

2026-07-09 health informatics 10.64898/2026.07.06.26357353 medRxiv
Top 0.1%
21.2%
Show abstract

Background: Real-world data (RWD) from different countries are increasingly used to support regulatory, health technology assessment (HTA), and population-level evidence generation. However, cross-country analyses are challenged by differences in data provenance, healthcare systems, coding practices, completeness, and clinical workflows. The Observational Medical Outcomes Partnership (OMOP) common data model (CDM) is widely used to harmonise heterogeneous RWD sources, but its ability to improve comparability of downstream epidemiological analyses relative to native source data across countries requires empirical evaluation. Methods: We examined RWD from Denmark, Finland and Portugal in their ability to capture epidemiology of female breast cancer (BC) and amyotrophic lateral sclerosis (ALS), exemplifying, respectively, a common disease with established treatment modalities and high survival and a rare fatal disease with scarce treatment options. To enable head-to-head comparison on a semantic level, data were mapped to the OMOP CDM. Data in the native format were used for comparison. In a downstream analysis, we examined disease epidemiology, patient characteristics, treatment, and survival. Results: OMOP conversion enabled a common analytical framework across countries and supported semantically aligned comparisons of key epidemiological and clinical variables. However, cross-country comparability was influenced by differences in data provenance, population coverage, coding practices, availability of clinical details, treatment capture, and healthcare-system-specific workflows. Iterative comparison with native data and external clinical evidence was necessary to identify mapping issues, assess information loss, and ensure high semantic fidelity of the converted data. Overall, OMOP-based estimates were highly consistent with native-data analyses and existing clinical expectations, but residual discrepancies reflected both source-data heterogeneity and decisions in the Extract, Transform, Load (ETL) workflow design. Conclusions: OMOP CDM conversion facilitates semantically meaningful cross-country analyses of RWD by mapping heterogeneous source data to a common structure and standardised vocabularies. However, CDM conversion does not eliminate heterogeneity in the underlying data-generating processes and cannot substitute for study-specific data quality and fitness-for-purpose assessment. Robust use of harmonised RWD for regulatory, HTA, or population-level evidence generation requires iterative benchmarking against native data, clinical expertise, and data-science expertise to support valid interpretation across countries.

20
A multimodal dataset for reconstructing common marmoset body-environment interactions in a 3D digital-twin framework

Iwata, K.; Kaneko, T.; Caeiro, C. C.; Thomas, D.; Koketsu, D.; Nambu, A.; Miyabe-Nishiwaki, T.; Hata, J.; Nakae, K.

2026-06-11 animal behavior and cognition 10.64898/2026.06.07.730757 medRxiv
Top 0.1%
18.9%
Show abstract

The common marmoset (Callithrix jacchus) is an important non-human primate model in neuroscience and biomedical research. However, existing 3D resources for this species have mainly focused on brain atlases or keypoint-based pose estimation, and reusable data resources that jointly describe the body surface, fur, articulated structure, and experimental environment remain limited. Here, we present a multimodal dataset designed to reconstruct body- environment interactions of common marmosets in three dimensions. The dataset includes a whole-body surface mesh derived from computed tomography (CT) images, fur representations based on photographic references, a rigged 3D model for pose-driven animation, synchronized behavioral videos from three individuals recorded for approximately 90 hours from eight view-points, 2D and 3D keypoint estimation data, 3D models of the experimental environment constructed from blueprint information, and rendered pseudo-egocentric views generated by integrating pose estimation results with the 3D body and environment models. Technical validation assessed the geometric agreement between the CT-derived mesh and the surface model, the accuracy of 2D and 3D keypoint estimation, the dimensional accuracy of the environment model, and the structural similarity between real and rendered images. This dataset provides a foundation for treating marmoset natural behavior not only as point trajectories but also as a three-dimensional phenomenon involving body shape and its spatial relationship with the environment, thereby enabling applications in behavioral analysis, visualization, synthetic-data generation, and future digital-twin studies.